A Discriminative Latent Model of Image Region and Object Tag Correspondence
نویسندگان
چکیده
We propose a discriminative latent model for annotating images with unaligned object-level textual annotations. Instead of using the bag-of-words image representation currently popular in the computer vision community, our model explicitly captures more intricate relationships underlying visual and textual information. In particular, we model the mapping that translates image regions to annotations. This mapping allows us to relate image regions to their corresponding annotation terms. We also model the overall scene label as latent information. This allows us to cluster test images. Our training data consist of images and their associated annotations. But we do not have access to the ground-truth regionto-annotation mapping or the overall scene label. We develop a novel variant of the latent SVM framework to model them as latent variables. Our experimental results demonstrate the effectiveness of the proposed model compared with other baseline methods.
منابع مشابه
Image Classification with Reconfigurable Spatial Structures
We propose a new latent variable model for scene recognition. Our approach represents a scene as a collection of region models (“parts”) arranged in a reconfigurable pattern. We partition an image into a pre-defined set of regions and use a latent variable to specify which region model is assigned to each image region. In our current implementation we use a bag of words representation to captur...
متن کاملشناسایی نوع و مدل وسیله نقلیه با استفاده از مجموعه بخشهای متمایزکننده
In fine-grained recognition, the main category of object is well known and the goal is to determine the subcategory or fine-grained category. Vehicle make and model recognition (VMMR) is a fine-grained classification problem. It includes several challenges like the large number of classes, substantial inner-class and small inter-class distance. VMMR can be utilized when license plate numbers ca...
متن کاملVinod Nair A thesis submitted in conformity with the requirements for the degree of Doctor of Philosophy
Visual Object Recognition Using Generative Models of Images Vinod Nair Doctor of Philosophy Graduate Department of Computer Science University of Toronto 2010 Visual object recognition is one of the key human capabilities that we would like machines to have. The problem is the following: given an image of an object (e.g. someone’s face), predict its label (e.g. that person’s name) from a set of...
متن کاملObject Recognition with Latent Conditional Random Fields
In this thesis we present a discriminative part-based approach for the recognition of object classes from unsegmented cluttered scenes. Objects are modelled as flexible constellations of parts conditioned on local observations. For each object class the probability of a given assignment of parts to local features is modelled by a Conditional Random Field (CRF). We propose an extension of the CR...
متن کاملA Discriminative Latent Model of Object Classes and Attributes
We present a discriminatively trained model for joint modelling of object class labels (e.g. “person”, “dog”, “chair”, etc.) and their visual attributes (e.g. “has head”, “furry”, “metal”, etc.). We treat attributes of an object as latent variables in our model and capture the correlations among attributes using an undirected graphical model built from training data. The advantage of our model ...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2010